Tag: AI factories

2 reviews

Selling infrastructure to inference providers: what a neocloud can offer without competing

Which inference providers should an AI-factory builder target, at which infrastructure layer, and what must it know about inference to sell them capacity without becoming their competitor?

The inference-provider market splits into open-weight GPU hosts (Together, Fireworks, DeepInfra, Baseten, Modal — they rent nearly all their capacity and compete on serving software) and closed-weight labs (OpenAI, Anthropic, Google, Meta, xAI — they rent at enormous scale via take-or-pay contracts but vertically integrate software and increasingly silicon). The evidence says a neocloud should target the open-weight hosts at the fleet/facility and control-plane layers — power, cooling, density, grid, orchestration that respects the customer's serving engine — and never compete at the serving-engine layer, which is the customer's moat. Firmus's public record (AI FactoryOS, HyperCube, Model-to-Grid, a Fireworks partnership) already points this way. Confidence is moderate: the market facts are industry-reported, and several claimed product names could not be verified in any public source.

Updated 23 Aug 202626 sources2022–2026Standard21 min read

inference providers · neocloud · GPU cloud · AI factories · LLM inference · take-or-pay · tokens per watt

The LLM compute market — training, fine-tuning, and inference

How is LLM compute demand, revenue, and profit split between training, fine-tuning, and inference, and what does the shift toward inference mean for AI factory builders?

Companies spent $37B on enterprise generative AI in 2025 and Gartner counts $2.59T of total AI spending for 2026, and the mix is shifting decisively from training to inference — inference overtakes training in AI-optimized cloud spend by 2026 and across the wider market by 2029. Fine-tuning is a real but small slice of that spend, agentic workloads are the fastest-growing inference category, and frontier training is multimodal-first even though public accounting of compute-by-modality barely exists. The evidence says an AI factory builder should not sell inference tokens against its own customers, but should become inference-workload-aware as infrastructure — disaggregated prefill and decode, SLO-aware scheduling, and tokens-per-watt are all sellable attributes.

Updated 19 Aug 202628 sources2019–2026Standard19 min read

LLM inference · training economics · fine-tuning · agentic AI · AI factories · GPU cloud